[ExecuTorch][llm] Fuse w1+w3 into single GEMM in quantized_moe_ffn#21124
[ExecuTorch][llm] Fuse w1+w3 into single GEMM in quantized_moe_ffn#21124digantdesai wants to merge 4 commits into
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21124
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 New Failure, 1 Unrelated FailureAs of commit ccdcc0b with merge base bfed808 ( NEW FAILURE - The following job has failed:
BROKEN TRUNK - The following job failed but were present on the merge base:👉 Rebase onto the `viable/strict` branch to avoid these failures
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
Stack from ghstack (oldest at bottom):
Fuse the up-projection (w1) and gate-projection (w3) into a single [2F, D] GEMM per expert. This halves the number of torchao activation quantizations per expert (from 2 to 1) and reduces total GEMM calls from 3 to 2 per active expert.
At AOT time, w1 and w3 are concatenated before packing: pack_fn(cat([w1, w3], dim=0)). At runtime, a single expert_linear_dispatch produces [m_e, 2F], then a fused swiglu_and_compact pass reads the interleaved h1/h3 and writes [m_e, F] for the w2 down-projection.
Schema changes from (packed_w1, packed_w3, packed_w2) to (packed_w13, packed_w2) — one fewer tensor arg (14 -> 13).
Differential Revision: D102799854